GRPO: Group Relative Policy Optimization

GRPO is a method introduced by DeepSeek Math, a variant of PPO that enhances mathematical reasoning abilities while concurrently optimizing the memory usage of PPO.

Date: 2026-09-08 Tue

Author: ArcaLunar